Papers with singing voice synthesis
Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and Accompaniment (2024.acl-long)
Copied to clipboard
Zhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang, Ruiqi Li, Fuming You, Zhou Zhao, Zhimeng Zhang
| Challenge: | Existing studies focus on singing voice synthesis and music generation independently. |
| Approach: | They propose a novel task called Text-to-Song synthesis which incorporates both vocal and accompaniment generation. |
| Outcome: | The proposed method can synthesize songs with comparable quality and style consistency. |
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on speech-to-singing voice conversion (STS) are limited by the scarcity of paired speech-song data and the suboptimal quality of outputs. |
| Approach: | They propose a self-supervised singing voice pre-training model that transforms a speech-to-singing voice into a paired singing voice. |
| Outcome: | The proposed model improves both STS and singing voice synthesis tasks by combining spoken language and a self-supervised singing voice pre-training model. |
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation (2025.findings-acl)
Copied to clipboard
Wenxiang Guo, Yu Zhang, Changhao Pan, Zhiyuan Zhu, Ruiqi Li, ZheTao Chen, Wenhao Xu, Fei Wu, Zhou Zhao
| Challenge: | Existing automated singing annotation (ASA) methods tackle isolated aspects of the annotation pipeline. |
| Approach: | They propose a framework that addresses transcription, alignment, and refined style annotations. |
| Outcome: | The proposed framework delivers comprehensive multi-level annotations encompassing: (1) precise phoneme-audio alignment, (2) robust note transcription and temporal localization, (3) expressive vocal technique identification, and (4) global stylistic characterization including emotion and pace. |
A Unified Feature Mixture Framework for Joint Speech and Singing Deepfake Detection (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for deepfake detection fail under speech-to-singing domain shift . a speech-retentive multi-domain fine-tuning strategy enables adaptation to singing . |
| Approach: | They propose a unified deepfake detector based on a multi-branch mixture-of-experts architecture that integrates three complementary feature views. |
| Outcome: | The proposed detector achieves 1.82% EER on CtrSVDD, compared to 37–62% for existing detectors . it can generalize to unseen generators and preserve strong speech performance . |